Skip to main content

Observability

Observability is the ability to answer questions about a running system that you did not anticipate when you built it. Monitoring tells you the interoperability layer is returning errors; observability tells you it is returning errors only for one facility's laboratory messages, and only since yesterday's terminology release.

In a health exchange, the operational question is usually which participant is broken right now — and the exchange is the only component positioned to know.


The three signals​

SignalAnswersCost
LogsWhat happened in this specific caseHigh volume, high storage
MetricsHow much, how often, how fast — aggregatedCheap, low cardinality only
TracesWhere the time went across servicesModerate; sampling usually required

A fourth, health-specific: audit records — who accessed which patient's data. These are not observability data. They contain personal health information, have legal retention requirements, and must be stored separately with their own access controls. Do not put AuditEvent data in the same store as application logs.


The cardinal rule​

No personal health data in telemetry.

Logs, metrics and traces are typically shipped to a monitoring system, retained loosely, and read by engineers with no clinical relationship to any patient. A log line containing a patient name or a diagnosis is a disclosure.

Practical rules:

  • Log identifiers, never names, addresses or clinical content
  • Prefer internal surrogate identifiers to national identifiers in logs
  • Never log request or response bodies for clinical payloads by default — and where an integration genuinely needs it for debugging, gate it behind an explicit, time-limited, audited setting
  • Never put a patient identifier in a metric label — it is both a privacy problem and a cardinality explosion
  • Scrub tokens, credentials and authorisation headers
  • Test for leakage: grep the log store for known synthetic patient names as a routine check

The OpenHIM transaction log is a deliberate exception: it persists bodies by design, which is why it must be treated as clinical data storage rather than as logging.


What to instrument in a health exchange​

Per integration channel​

  • Message volume, by source system and message type
  • Error rate, by error class — validation failure, identity unresolved, downstream timeout, authorisation denied
  • Latency distribution, at p50/p95/p99 — averages hide the cases that matter
  • Queue depth and queue age. Depth alone misses the slow drain; a queue with fifty items that are six hours old is a data-freshness incident
  • Retry and dead-letter counts

Per participant system​

  • Last successful message received. This is the single most useful signal in an exchange: a facility that stopped sending three days ago is invisible in aggregate error rates
  • Availability of the system's endpoint
  • Conformance failure rate, which spikes after either side upgrades

Business-level signals​

Technical health is not the same as the exchange working. Instrument outcomes:

  • Records successfully linked to a client registry identity, versus routed to the review queue
  • Review queue depth and age — see MPI
  • Terminology translation failures, by code and source
  • Reporting completeness by facility for the current period
  • Break-glass invocations

These are the metrics that tell you whether the architecture is delivering value. A dashboard showing 99.9% uptime and a six-week-old review queue is describing a failed exchange.


SLIs, SLOs and SLAs​

  • SLI — a measured indicator: "proportion of $match requests completing under 500 ms"
  • SLO — the internal target: "99% over 30 days"
  • SLA — the contractual commitment, with consequences; usually looser than the SLO

Set SLOs from clinical consequence, as with availability tiers. Patient lookup during registration has a tight latency SLO because a queue forms at the desk. Overnight HMIS submission does not.

Error budgets work well in this setting: if the SLO is 99.9% over 30 days, about 43 minutes of failure is acceptable, and that budget is what pays for upgrades and change. When it is exhausted, change stops until reliability is restored. This gives a health ministry a defensible, non-arbitrary rule for slowing down deployments.


Health checks​

Three kinds, and conflating them causes outages:

  • Liveness — is the process alive? If not, restart it.
  • Readiness — can it serve traffic now? Remove from the load balancer if not.
  • Dependency health — can it reach the registry, the terminology service, the database?

The failure to avoid: a readiness check that fails when a downstream dependency is unavailable. The orchestrator then removes every instance from service, and a degraded dependency becomes a total outage. Report dependency health as a distinct signal; keep readiness about the service itself.


Tooling​

ToolRoleLicence
OpenTelemetryVendor-neutral instrumentation for traces, metrics and logs — the right default, because it decouples instrumentation from backend choiceApache 2.0
PrometheusMetrics collection and alertingApache 2.0
GrafanaDashboards across metrics, logs and tracesAGPL
LokiLog aggregation, indexed by labelAGPL
OpenSearch / ElasticsearchLog search and analyticsApache 2.0 / SSPL
Jaeger / TempoDistributed tracing backendsApache 2.0 / AGPL
AlertmanagerRouting, grouping and silencing alertsApache 2.0
Uptime Kuma / Blackbox exporterExternal endpoint checksMIT / Apache 2.0

Instrument with OpenTelemetry regardless of backend choice. It is the decision that is cheapest to make now and most expensive to retrofit.


Alerting discipline​

An alert that does not require a human to act on it immediately should not page anyone. Health operations teams are small; alert fatigue is the reason real alerts get missed.

For each alert define: what is broken, what the clinical or data impact is, what the responder should do first, and how to confirm resolution. If any of those cannot be written, the alert is a dashboard item.

Alert on symptoms, not causes: "laboratory results have not reached the SHR for 30 minutes" is actionable; "CPU is at 80%" is not.


References​